How we
solved prefill.
Prefill hit its compute wall. Memory optimisations won't get past it. Photonics will.
Inference has two steps.
The ratio of input to output tokens is 15:1, and growing. 93% of tokens are prefill. The only way to speed up prefill is to speed up the core itself.
Processes the prompt.
Generates the KV cache.
- OperationMATRIX × MATRIX
- BottleneckCOMPUTE
Generates output tokens, one at a time.
- OperationVECTOR × MATRIX
- BottleneckMEMORY BW
Coding is eating
the AI market.
50% of all tokens now go to programming. Share increased 4.5× in six months. AI revenue is growing 10×/yr. Coding workloads are almost entirely prefill — 99% of agentic coding tokens are input tokens. The workload the market is shifting toward is exactly the one OPUs dominate.
The industry is solving decode.
Photonic compute is the only path through prefill.
Decode has a roadmap: MHA → GQA → MLP → LRKV → QSVD. HBM to distributed SRAM. Memory bandwidth is the bottleneck, and the industry is making rapid progress there. Prefill is compute-bound. Not solvable with memory optimizations. The only way forward is to speed up the core itself — and transistor scaling has slowed. This is what the OPU was built for.
Three locked doors.
Each one has now opened.
There was no market for 4-bit GEMM.
There were no small optical transistors.
Moore's Law was still delivering.
“There are no small optical transistors.”
Neurophos
shrank them
10,000×.
Speed = clock
× ops per clock.
56 GHz clock — 30× NVIDIA Blackwell. 8× operations per clock as arrays reach 3k × 3k through quadratic compute scaling. In metaphor: a 56,000-round clip fired at 56 GHz, each bullet several million ops — reload at a MHz.
420× faster than
NVIDIA Blackwell.
Photonic cores run at 56 GHz — 30× faster than Blackwell. 8× more operations per clock, scaling quadratically with array size. Energy efficiency compounds as we scale — optical cores consume power only for I/O.
The entire industry
fits in the bottom corner.
Speed (TOPS/mm²) on the vertical. Efficiency (TOPS/W) on the horizontal. Everyone else clusters near the origin. Neurophos OPUs plot exponentially off-chart.
| METRIC | NEUROPHOS OPU · MANWE | NVIDIA B200 |
|---|---|---|
| PEAK SPEED | 4.2 ExaOPS | ~10 PetaOPS |
| ENERGY EFFICIENCY | 1,400 TOPS/W | ~9 TOPS/W |
| COMPUTE DENSITY | ~2,800 TOPS/mm² | ~5 TOPS/mm² |
| CLOCK SPEED | 56 GHz | 1.84 GHz |
| POWER DRAW | 3,000 W | 1,000 W |
| PHYSICS | Photonic tensor cores | Electronic transistors |
| SCALING LAW | Exponential (optical) | Slowing (Moore's Law) |
Now touch it.
Six views of the T100 OPU. Everything from the wafer down to a single waveguide.